Papers by Severino Da Dalt
FLOR: On the Effectiveness of Language Adaptation (2024.lrec-main)
Copied to clipboard
Severino Da Dalt, Joan Llop, Irene Baucells, Marc Pamies, Yishi Xu, Aitor Gonzalez-Agirre, Marta Villegas
| Challenge: | Large language models have amply proven their capabilities, but low- and mid-resource languages do not have access to the necessary means to train such models from scratch. |
| Approach: | They use a 26B tokens corpus to further pre-train BLOOM, giving rise to FLOR models. |
| Outcome: | The proposed model achieves consistent gains across Catalan and Spanish tasks. |
A CURATEd CATalog: Rethinking the Extraction of Pretraining Corpora for Mid-Resourced Languages (2024.lrec-main)
Copied to clipboard
Jorge Palomar-Giner, Jose Javier Saiz, Ferran Espuña, Mario Mina, Severino Da Dalt, Joan Llop, Malte Ostendorff, Pedro Ortiz Suarez, Georg Rehm, Aitor Gonzalez-Agirre, Marta Villegas
| Challenge: | CATalog 1.0 is the largest text corpus in Catalan to date . CURATE is a pipeline that can be parallelizable to run in high performance clusters . |
| Approach: | They propose a data pipeline that uses binary filters to filter documents based on text quality . they optimised the pipeline to run in high performance clusters . |
| Outcome: | The proposed pipeline is optimized for high performance cluster environments and runs in high performance. |